Papers with expected validation performance
Show Your Work: Improved Reporting of Experimental Results (D19-1)
Copied to clipboard
| Challenge: | Current practice is to train multiple instantiations of each, choose the best model of each type, and compare their performance on held-out test data. |
| Approach: | They propose to measure expected validation accuracy as a function of computation budget . authors find comparisons where authors would have reached different conclusions if they had used more computation . |
| Outcome: | The proposed method shows that test-set performance scores alone are insufficient for drawing accurate conclusions about which model performs best. |